Papers with dense representations

19 papers
Pretrained Transformers for Text Ranking: BERT and Beyond (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial provides an overview of text ranking using neural network architectures known as transformers.
Approach: This tutorial provides an overview of text ranking with neural network architectures known as transformers.
Outcome: This tutorial provides an overview of text ranking with neural network architectures known as transformers.
Extracting Text Representations for Terms and Phrases in Technical Domains (2023.acl-industry)

Copied to clipboard

Challenge: Large pre-trained language models are extensively used in modern NLP systems.
Approach: They propose an unsupervised approach to encoding using character-based models and pre-trained sentence encoders to reconstruct large pre-trained embedding matrices.
Outcome: The proposed approach matches the quality of sentence encoders in technical domains and is 5 times smaller and up to 10 times faster on high-end GPUs.
GNN-encoder: Learning a Dual-encoder Architecture via Graph Neural Networks for Dense Passage Retrieval (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to perform large-scale query-passage retrieval are term-based, but they lose interaction between query-pastage pairs.
Approach: They propose to fuse query (passage) information into query representations via graph neural networks that are constructed by queries and their top retrieved passages.
Outcome: The proposed model outperforms existing models on MSMARCO, Natural Questions and TriviaQA datasets and achieves the new state-of-the-art on these datasets.
Contextualized Query Embeddings for Conversational Search (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to conversational search use multiple inference pipelines that require long inference times . despite their effectiveness, such a pipeline often includes multiple neural models that require longer inference time.
Approach: They propose to integrate conversational query reformulation directly into a dense retrieval model . they use a dataset with pseudo-relevance labels to overcome the lack of training data .
Outcome: The proposed model rewrites conversational queries as dense representations in conversational search and open-domain question answering datasets.
The Curse of Dense Low-Dimensional Information Retrieval for Large Index Sizes (2021.acl-short)

Copied to clipboard

Challenge: Existing studies have shown that dense representations outperform sparse representations with large index sizes.
Approach: They propose to use dense low-dimensional representations to retrieve relevant documents . they show performance decreases quicker for increasing index sizes than for sparse representations .
Outcome: The proposed representations outperform sparse representations with large index sizes.
Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval (2021.emnlp-main)

Copied to clipboard

Challenge: Recent approaches to information retrieval (IR) and natural language processing (NLP) use contextual language models, which can improve both synonymy and polysemy problems associated with words.
Approach: They propose an ultra-high dimensional representation scheme equipped with directly controllable sparsity and a bucketing method where embeddings from multiple layers of BERT are selected/merged to represent diverse linguistic aspects.
Outcome: The proposed representation scheme outperforms sparse models with MS MARCO and TREC CAR, and shows that it is highly efficient for storage and search.
Explaining Relationships Between Scientific Documents (2021.acl-long)

Copied to clipboard

Challenge: Existing approaches to explain relationships between scientific documents using natural language text can be useful for research efficiency.
Approach: They propose a task of explaining relationships between scientific documents using natural language text.
Outcome: The proposed models can be automated and humanely evaluated.
Learning Target-Specific Representations of Financial News Documents For Cumulative Abnormal Return Prediction (C18-1)

Copied to clipboard

Challenge: Recent work considers learning dense representations for news titles and abstracts . text representations can address the sparsity of discrete indicators in statistical models .
Approach: They propose to use news abstracts to combine the most informative sentences in news content to learn dense representations for text elements.
Outcome: The proposed model can be used to estimate abnormal returns of companies when compared to titles and abstracts.
Transformation of Dense and Sparse Text Representations (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to NLP to leverage sparsity have been limited due to the gap with dense representations.
Approach: They propose a Semantic Transformation method to bridge dense and sparse spaces and propose supervised NLP tasks to use both spaces.
Outcome: Experiments with classification tasks and natural language inference tasks show that the proposed method is effective.
Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval (2021.naacl-main)

Copied to clipboard

Challenge: Current methods for complex question answering use structured knowledge and unstructured text.
Approach: They propose a multi-step retrieval approach that iteratively forms an evidence chain through beam search in dense representations.
Outcome: The proposed method is competitive to state-of-the-art systems without using semi-structured information.
Unifying Multimodal Retrieval via Document Screenshot Embedding (2024.emnlp-main)

Copied to clipboard

Challenge: Existing document retrieval pipelines require document parsing and content extraction to prepare input for indexing.
Approach: They propose a retrieval paradigm that regards document screenshots as a unified input format . they leverage a large vision-language model to directly encode document screenshot into dense representations .
Outcome: The proposed method outperforms existing retrieval pipelines in a text-intensive context.
Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval (2021.acl-long)

Copied to clipboard

Challenge: Existing retrieval models based on dense representations show better performance than sparse representations.
Approach: They propose a method to mimic the queries to each of the documents by an iterative clustering process and represent the documents using multiple pseudo queries.
Outcome: The proposed model achieves state-of-the-art results on a large dataset while remaining high efficiency.
Natural Logic-guided Autoregressive Multi-hop Document Retrieval for Fact Verification (2022.emnlp-main)

Copied to clipboard

Challenge: Recent evidence retrieval approaches rely on heuristics and assume hyperlinks between documents.
Approach: They propose a retrieval method that combines a retriever and a proof system that reranks documents and reorders them .
Outcome: The proposed method exceeds or is on par with the current state-of-the-art on FEVER, HoVer and FEVEROUS-S while using 5 to 10 times less memory than competing systems.
Learning Dense Representations of Phrases at Scale (2021.acl-long)

Copied to clipboard

Challenge: Existing phrase retrieval models rely on sparse representations and still underperform retriever-reader approaches.
Approach: They propose a method to learn phrase representations from reading comprehension tasks using negative sampling methods.
Outcome: The proposed model improves over previous models by 15%-25% absolute accuracy and matches the performance of state-of-the-art retrieval models.
Dense Passage Retrieval for Open-Domain Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Open-domain question answering relies on efficient passage retrieval to select candidate contexts.
Approach: They propose a dual-encoder framework that can be implemented to retrieve passages from a small number of questions and passages.
Outcome: The proposed system outperforms a strong Lucene-BM25 system in top-20 passage retrieval accuracy on multiple open-domain QA benchmarks.
MIST: Mutual Information Maximization for Short Text Clustering (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for clustering short texts are inadequate due to the limited amount of information provided by each text sample.
Approach: They propose a Mutual Information Maximization Framework for Short Text Clustering which maximizes mutual information between representations on sequence and token levels.
Outcome: The proposed framework outperforms the state-of-the-art method in terms of Accuracy or Normalized Mutual Information in most cases.
STAIR: Learning Sparse Text and Image Representation in Grounded Tokens (2023.emnlp-main)

Copied to clipboard

Challenge: State-of-the-art contrastive learning models like CLIP and ALIGN are less interpretable and suffer from inferior accuracy than dense representations.
Approach: They extend CLIP and ALIGN models to build a sparse semantic representation that is interpretable and easy to integrate with existing retrieval systems.
Outcome: The proposed model outperforms CLIP and ALIGN models on image and text retrieval tasks with a 4.9% and +4.3% improvement on COCO-5k textimage and imagetext retrieval respectively.
ConvX: A Lightweight Converter to Bridge Indexed Dense Representations and Large Language Models for Retrieval-Augmented Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing RAG pipelines suffer from critical efficiency limitations due to their complexity and complexity.
Approach: They propose a compression-based RAG framework that directly leverages indexed dense representations produced by a retriever, substituting to long text contexts.
Outcome: Empirical results show that the proposed model achieves competitive performances compared to the state-of-the-art model that uses a large ad-hoc context compressor while offering substantially improved inference efficiency.
Decoupled Reasoning with Implicit Fact Tokens (DRIFT): A Dual-Model Framework for Efficient Long-Context Inference (2026.findings-acl)

Copied to clipboard

Challenge: Existing solutions to integrate extensive, dynamic knowledge into Large Language Models (LLMs) are constrained by finite context windows, retriever noise, or the risk of catastrophic forgetting.
Approach: They propose a dual-model architecture that explicitly decouples knowledge extraction from the reasoning process by compressing document chunks into implicit fact tokens conditioned on the query.
Outcome: The proposed architecture significantly outperforms strong baselines among comparably sized models on long-context tasks while maintaining inference accuracy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations